Conversation
4d9994e to
49ad3ac
Compare
fd089c2 to
9e1b350
Compare
9e1b350 to
520c0f9
Compare
|
Found two issues in the text verification:
|
hsliuustc0106
left a comment
There was a problem hiding this comment.
One additional P2 finding in text-field readback.
| current = [e for e in self._require_snapshot().elements if self._locator(e) == locator] | ||
| if result.get("effect") != "confirmed": | ||
| raise DriverError(f"text verification failed: driver effect {result.get('effect', 'missing')}") | ||
| value = current[0].value if len(current) == 1 else None |
There was a problem hiding this comment.
[P2] Preserve text-field identity through readback
Please reject known-ambiguous text locators before offering or executing the write, or retain stable structural identity for readback. Two fields sharing role and label without AXIdentifier get distinct offered keys, but _locator collapses them afterward. A successful write to type:Body#1 therefore always raises because len(current) == 2, leaving the document modified and the episode failed.
There was a problem hiding this comment.
Fixed in 041a71e. Ambiguous text locators are excluded before input. Same-label fields with distinct identifiers remain usable.
Added tests for duplicate labels, duplicate identifiers, disabled duplicates, and readback with distinct identifiers. Local full suite: 522 passed, 41 skipped. GitHub CI still needs maintainer approval.
|
@linear3735 Fixed both cases in 17fa36e. Success now requires verified text. A confirmed same-text selection replacement is accepted after fresh readback, without requiring the payload count to increase. Added regression tests for both cases and kept the partial-input and unverified-input checks. Full suite at 041a71e: 522 passed, 41 skipped. Ruff, ty, build, and CLI/MCP smoke passed. GitHub CI is waiting for maintainer approval. |
|
@QianCyrus could you add a Demo / evidence section following the self-review guidance? I see the verification fixes and regression tests in 041a71e; please link an inspectable artifact for the input-to-save workflow. Show the supplied text, field readback, Save action and independently verified fresh file, with the Laya-selected run separate from the fixed-plan Chinese/multiline check. Please include current-head regression output for Save-before-input, same-text selection replacement and ambiguous-field rejection; these can use the existing fake-driver tests rather than another native/model run. Link the episode/logs with the exact commit, checkpoint, commands and Mac/driver settings. Label any older native results with their actual source revision, retain failures/interventions, and keep load time separate from episode time. A recording, screenshots with logs, or a terminal trace is fine. Please redact private data and state media reuse/attribution terms. |
|
@hsliuustc0106 Added the evidence in 1ffac8b and linked it from the PR. Reran Laya (3/3) and the separate Chinese/multiline fixed plan (1/1). The trace keeps the early Save clicks and includes field readback, fresh-file checks, pinned versions and the requested regression output. Local suite: 522 passed, 41 skipped; lint, types, build and smoke passed. CI still needs maintainer approval. |
|
@QianCyrus Please try the text-entry desktop agent with system1-omni serving Laya, and attach a short video under this PR. The current learned-model evidence uses in-process Laya 0.3.5 on MPS. System1-Omni's Laya MPS/CPU worker is supported; #35 supplies the The video should show the supplied text, actual model-selected input/Save actions, field readback and independently verified fresh file. Include the agents/engine commits, checkpoint, command and hardware, and link the episode/logs. Identify the separate fixed-plan Chinese/multiline demonstration when showing it. |
hsliuustc0106
left a comment
There was a problem hiding this comment.
Re-reviewed 1ffac8b. Two P2 findings remain below: completion can still accept a Saved status from before the text input when --verify-file is omitted, and the new evidence document fails the locked Ruff formatter.
The same-text selection replacement and ambiguous-field rejection fixes are verified. I also checked the native transcript's source revision, separation of Laya and fixed-plan runs, retained early Save attempts, independent file checks, and recorder hash.
Validation on Linux / Python 3.12.13: 69 targeted tests passed; ruff check, ty check, uv lock --check --offline, and CLI/MCP smoke passed. ruff format --check . fails on text-evidence.md. Native desktop execution, real Laya inference, and the full test suite were not rerun. Current-head upstream CI is still awaiting approval: https://github.com/ThinkFlowLab/system1-agents/actions/runs/37272336366.
| def score(self) -> float: | ||
| if self._text and not self._typed: | ||
| return 0.0 | ||
| return 1.0 if self._snapshot is not None and self._done_when(self._snapshot) else 0.0 |
There was a problem hiding this comment.
[P2] Invalidate completion evidence that predates text input
With --text hello --expect Saved and --plan 'Save,type:Body,Save', without --verify-file, the actual rule-agent loop stops after Save and type:Body. I reproduced this using the existing FakeDocument as the driver boundary: mean_score=1.0, errors=0, field readback='hello', but saved_file=''. The final Save never executes. The native fixture likewise leaves its Saved status unchanged when Body is edited.
Once _typed becomes True, this check accepts the status left by the earlier Save, so the pending-input guard only postpones the false success. Adding --verify-file makes the final Save execute and produces 'hello' on disk. Please invalidate completion evidence that predates the verified text write and add a full-agent regression that asserts the final Save occurs; the current Save-before-input test calls that final step manually and does not check whether the agent would already have stopped.
There was a problem hiding this comment.
Fixed in cbeeafb. Text input now invalidates a pre-existing completion result. The agent must observe a new result or repeat the action that produced it; an unrelated click cannot clear the guard.
Added full-agent regressions for insert and replace, with early Save, missing final Save and unrelated clicks. They check the saved file independently. Local suite: 524 passed, 41 skipped; lint, types and smoke passed. The new CI run needs maintainer approval.
| data = capture.image.data | ||
| digest = hashlib.sha256(data).hexdigest() | ||
| (root / f"{digest}.png").write_bytes(data) | ||
| fields["capture"] = {"id": capture.capture_id, "width": capture.width, "height": capture.height, "sha256": digest} |
There was a problem hiding this comment.
[P2] Format the recorder code block before submitting the evidence
The locked Ruff 0.16.8 checks Python code blocks in Markdown. At this head, ruff format --check . exits 1 on this dictionary and the record(...) call at lines 315-316, so the documented local format pass does not cover the newly added evidence file. Please run the formatter on this document and rerun the current-head check. Since that changes the recorder source bytes, also update its recorded SHA-256.
There was a problem hiding this comment.
Fixed in cbeeafb. Formatted the recorder block and updated its SHA-256. The original run hash and source link are retained separately; the parsed Python syntax trees match. ruff format --check . now passes with the locked Ruff 0.16.8, including the evidence document.
|
@hsliuustc0106 Checked the serving dependency: #35 is still open, and this branch does not yet have The served run and its video remain pending #35. The existing evidence is still labelled as in-process Laya, with the Chinese/multiline fixed plan separate. |
What this enables
The desktop agent can now enter supplied text into a macOS application, click Save, and check that the file contains the expected text.
It can insert text at the current selection or replace a field's entire contents, including Chinese and multiple lines. After input, it reads the field again to confirm the result. An optional file check requires a fresh write and matching content before the task succeeds.
The implementation extends the existing desktop CLI, window environment, and Cua Driver adapter. Part of #15; RFC #26.
Demo / evidence
Input-to-save evidence: terminal traces, reproduction commands and regressions
Fresh runs on 2026-10-05, using source
041a71e, macOS 26.2 arm64 and Cua Driver 0.30.1:The artifact includes supplied text, field readback, Save actions and independent fresh-file checks for every trial. It retains Laya's two early Save clicks per episode, which did not count as completion. Model-selected runs, fixed-plan execution and fake-driver regressions are separate.
Exact source/checkpoint revisions, Mac/driver settings, commands, load time, warnings and reuse terms are documented. Episode time excludes loading and fixture reset. These fresh traces replace the earlier native figures whose original artifacts were unavailable.
Serving/video follow-up: The requested System1-Omni run and video are pending the
--model laya-servedclient in #35, which is still open. The native evidence above uses in-process Laya; this branch does not yet have the served client.Tests
Added
test_desktop_actions.pyfor text entry, replacement, readback, and saved-file verification.The full-agent regressions also cover Save-before-input without
--verify-file, for both insert and replace. They verify that the final Save executes and the file contains the text; omitting Save or clicking an unrelated button cannot report success.Extended
test_desktop_driver.py,test_agents_desktop.py, andtest_desktop_env.pyfor window binding, native field identifiers, CLI configuration, dry runs, and stale or refused actions.Core + dev suite: 524 passed, 41 skipped. Ruff, ty, lock check, package build, shell syntax, and CLI/MCP smoke passed.
GitHub CI for the current head is
action_required, awaiting maintainer approval; no upstream test job has run.How to run
Requires macOS,
uv, the Swift command-line tools, and Cua Driver with Accessibility and Screen Recording permissions. These commands use the signed driver at/Applications/CuaDriver.app/Contents/MacOS/cua-driver. Run from the current project checkout. Close any older S1A document fixture before starting.Run a fixed input-and-save sequence to check execution:
To let local Laya select the actions instead, run: